Value in Health
○ Elsevier BV
Preprints posted in the last 90 days, ranked by how well they match Value in Health's content profile, based on 11 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.
Brodsky, S.; Matlin, O.
Show abstract
Improving primary care is a long-standing strategy to constrain health care spending. Yet, evaluations of primary care models focused on payment reform have shown minimal effects on total cost of care. We report the results from a large-scale, real-world evaluation of an advanced primary care model that restructures access through same-day and next-day appointments, on-demand video visits, asynchronous clinician messaging, and extended hours. Using a stacked-cohort difference-in-differences design with entropy balancing and inverse probability of censoring weighting, we analyzed multi-payer claims covering April 2022 through March 2025. Advanced primary care use was associated with an 8.6% reduction in total cost of care (-$729 per patient per year; P = 0.004), driven by lower specialist cost (-$939/year; P < 0.001) and, to a lesser degree, by reductions in inpatient (-$134/year; P < 0.001), urgent care (-$70/year; P < 0.001), and emergency department cost (-$16/year; P = 0.02), partially offset by higher primary care cost (+$350/year; P < 0.001). The specialist reduction was concentrated in knowledge-based consultative encounters (-$663/year; P < 0.001), while procedural specialist cost was largely unchanged (-$276/year; P = 0.09). Cost differences emerged in the first post-index month. These findings suggest that advanced primary care may reduce total health care spending, with observed savings driven primarily by lower spending on consultative specialty care.
Creeden, J.; Olivecrona, M.; Soriano, A.
Show abstract
Background. Tertiary clinical genomics reports condense layered molecular findings into documents that treating oncologists must read, translate, and act upon; manual summarisation of these reports is time-consuming and variable. Tools that assist summarisation and translation into local languages are emerging, yet the field lacks an agreed methodology for evaluating such tools before any downstream clinical use. The appropriate first endpoint is fidelity of the generated summary to its source report, assessed by qualified human raters under blinded scoring, not downstream variant classification. Methods. QNOMX-VHIR-CPSP-001 Phase 1 is a single-site, non-interventional clinical performance study conducted at Vall d'Hebron Institut de Recerca (VHIR) under ISO 20916:2019 as a Clinical Performance Study Protocol. De-identified tertiary cancer genomics reports from pediatric oncology cases are summarised by the AI-assisted summarisation system under evaluation and, in parallel, by the standard manual workflow. Qualified raters score both summary types against the source genomics report using the Quality Summary Index (QSI), a six-dimension, five-point rubric adapted from the Provider Documentation Summarization Quality Instrument, under a blinded, counterbalanced, two-period crossover with a minimum fourteen-day washout. Two co-primary composite endpoints, content and presentation, are analysed for non-inferiority under a Bayesian hierarchical model, with a frequentist linear mixed model as the convergence check. Inter-rater reliability is reported as Krippendorff's ; a Monte-Carlo power analysis of the fixed clustered design is pre-specified. Discussion. The design isolates summarisation quality from clinical decision-making by scoring both summary types against the same source report under blinding, counterbalancing, and a fourteen-day washout. Conclusion. The QSI rubric, the counterbalanced crossover, and the pre-specified Bayesian primary with frequentist convergence check define a replicable protocol for early-stage evaluation of AI-assisted summarisation in tertiary genomics reporting; observed variance components will inform sample-size determination for Phase 2.
Reza, L.; Arbai, Z.; Ward, H.; Payne, L.; Kinross, J.; Patel, V.
Show abstract
Background Virtual hospital (VH) pathways support early discharge through remote monitoring, but limited evidence has hindered implementation in colorectal surgery. This study aimed to define patient- and carer-relevant outcomes and experiences of VH following colorectal surgery. Methodology A patient and public involvement and engagement (PPIE) consultation was conducted with 8 participants (7 patients, 1 carer; 4 women, 4 men) who had experienced VH following bowel resection at a high-volume robotic unit. Purposive sampling ensured that 50% of participants had experienced readmission. The 90-minute session was delivered via Microsoft Teams. Data were analysed using reflexive thematic analysis. Results Seven themes were identified: readmission, remote monitoring, carer burden, recovery, equity, readiness for discharge, and information delivery. Patients supported early discharge when remote monitoring enabled timely detection of complications and readmission pathways were efficient. Readmission was not perceived as failure but as appropriate escalation. Dissatisfaction with readmission was related to delays in emergency care. Remote monitoring provided psychological safety, with patients feeling held at home. Carers assumed substantial, often unrecognised, quasi-clinical roles. Recovery was defined by return to function rather than length of stay. Equity concerns were evident, with VH favouring those with adequate support at home, digital literacy, and language proficiency. Discharge readiness was both clinical and psychological. Information delivery at discharge was often poorly retained and requires reinforcement preoperatively at every encounter with patients and carers. Conclusions VH pathways are acceptable and valued. Readmission is a marker of system responsiveness rather than failure of early discharge on VH. Psychological preparedness, carer support, and equitable access are critical to successful and scalable implementation of early discharge using a virtual hospital.
Hamdan, M.; Harati, A.; Al-Bakheet, A.; Fuetterer, I.; Alshaer, I.
Show abstract
Objective: To evaluate decision concordance between commercially available multimodal large language models (LLMs), resident doctors, and senior-surgeon ground truth for surgical indication and spinal level in degenerative lumbar spine disease. Methods: We retrospectively analyzed 147 consecutive patients. Each case included clinical documentation and MRI presented as two composite PNG images. Two resident doctors and three multimodal LLMs (GPT 5.5, Claude Sonnet 4.6, Gemini 3.1 Pro) independently assessed operative versus conservative management and, if operative, the surgical level. Analyses used Cochran's Q, McNemar tests with Holm correction, and Bayesian methods. Results: LLMs achieved higher therapy-decision accuracy (66.0%-68.0%; 97-100/147) than residents (54.4%; 80/147) but over-recommended surgery. Conditional level accuracy when surgery was correctly indicated was 71.4% (20/28) for residents versus 33.3%-41.1% for LLMs. Conclusion: Off-the-shelf multimodal LLMs approximate human performance for binary surgical indication but remain inferior for precise level localization. These results establish a practice-relevant baseline of spatial reasoning limitations for tools already used by patients and junior doctors.
Girdwood, S.; Marban-Castro, E.; Muhwava, L.; de Beer, J. C.; Haldane, C.; Dave, J. A.; Carrihill, M.; Karsas, M.; Rheeder, P.
Show abstract
Background: Continuous glucose monitoring (CGM) improves glycaemic control in people with type 1 diabetes (T1D), but high costs limit uptake in low- and middle-income (LMIC) countries. Evidence on the cost-effectiveness of CGM is limited in LMIC settings. Objective: To evaluate the short- and long-term cost-effectiveness of intermittently-scanned continuous CGM (cCGM) and intermittently-scanned periodic CGM (pCGM) (one sensor every three months), versus standard self-monitoring of blood glucose (SMBG) in the South African public-sector. Methods: A three-arm randomised controlled trial (ACCEDE) was conducted among T1D individuals with HbA1c [≥]10% in South Africa. A within-trial cost-effectiveness analysis was conducted from a partial societal perspective over 9-months, using resource use data and QALYs derived from EQ-5D. A cost-utility analysis using a Markov microsimulation model was conducted for two populations: total T1D, and youth (<20 years). Results: Within-trial analysis showed no statistically significant differences in QALYs or HbA1c between arms. Costs were highest for cCGM (USD 1,504), followed by pCGM (USD 742) and SMBG (USD 467). CGM strategies were dominated in the within-trial analysis. In contrast, long-term modelling showed that CGM was more effective than SMBG and was cost-effective for youth when used periodically. cCGM delivered additional QALYs at a higher cost (ICERs USD 15,259-30,852/QALY) and was only potentially cost-effective in youth when sensor prices were reduced by >45%. Conclusions: While CGM was not cost-effective in the short term, modelling suggests pCGM use may offer value-for-money under specific assumptions and in select populations, highlighting the need for further evidence on long-term effectiveness, engagement, and pricing.
Leonhardt, C.; Birrer, D.; Stauffer, M. F.; Toti, J. M. A.; Gallagher, I. J.; Skipworth, R. J. E.; Laird, B.; Kuemmerli, C.
Show abstract
Importance Non-inferiority trials are becoming increasingly popular in abdominal surgery. The non- inferiority margin is critical in the interpretation and conclusion of these trials. Objective This systematic review aims to assess the methodological and reporting quality of non- inferiority randomized controlled trials in abdominal surgery. Evidence Review Non-inferiority trials were systematically identified by searching Ovid Medline, Embase and the CENTRAL databases from 2006 until December 2025. Randomized controlled trials in adult patients with any type of abdominal surgical intervention in at least one trial arm and a sample size greater than or equal to 100 were eligible for inclusion. The primary outcome was the definition of the non- inferiority margin. Secondary outcomes were the reporting of the non-inferiority margin, the robustness of its estimation, the uncertainty of the point estimate and the adequacy of conclusions. Findings A total of 11 045 trials were identified, of which 101 were eligible, enrolling 44 370 patients. Most trials provided a rationale for the non-inferiority design, while six (5.9%) trials did not. Previous literature was commonly used (n=56; 55.4%), but the non-inferiority margin was most often based on a clinical fixed margin or on historical comparison of the treatment and the active comparator. Based on the margin, investigators tolerated substantially worse outcomes of the treatment compared to the comparator. Conclusions were appropriate based on the confidence interval and the predefined non- inferiority margin in 88 (87.1%) of trials. The clinical judgement of the conclusion was overall adequate. Confidence interval estimations were reported in 16 (15.8%) of trials. Simulation studies were limited by the reporting quality. Conclusions and Relevance Clinical fixed margins are commonly used in abdominal surgery non-inferiority randomized controlled trials, however, substantial shortcomings in reporting limit the interpretability and reproduction of study findings. Based on the findings of this study, guidance on surgical- specific non-inferiority margin definitions is needed.
Howran, J.; Sharma, A.; Andrews, K.; Pople McCord, D.; Salim, S. K.; Janka, D.; Ynoe Moraes, F.; Goldie, K.; Babiolakis, C.; Alkins, R.; Taslimi, S.; Pasarikovski, C.; Ebinu, J.; Cook, D. J.; Levy, R.; Purzner, J.; Purzner, T.
Show abstract
Objective: To evaluate whether entrepreneurial methodologies applied to system-wide healthcare redesign were associated with improved survival, care timeliness, and rural-urban equity among patients with glioblastoma. Design: Non-randomized pre-post cohort study Setting: Tertiary neuro-oncology centre in Ontario, Canada Participants: Adults aged >18 years with histologically confirmed glioblastoma who underwent surgical resection between January 1, 2018, and February 28, 2025 Interventions: Implementation of the Integrative Brain Tumor Program (IBTP), a system-wide care intervention grounded in user-defined priorities and developed using a design thinking approach (empathize, define, ideate, prototype, test) integrated with operational frameworks adapted from early-stage innovation. System change was treated as a deliberate, deployable intervention that could be designed, launched, iteratively refined, and evaluated. Coordinated changes were embedded across healthcare services within existing infrastructure and resource constraints through a single centralized nurse navigator who standardized referrals, patient education, and real-time care coordination. Main Outcomes and Measures: Primary outcomes were one-year overall survival and time to postoperative MRI completion and radiotherapy initiation. Secondary outcomes assessed rural-urban equity in these measures. Associations were evaluated using Cox proportional hazards and Fine-Gray competing-risk models adjusted for age, sex, rurality, MGMT promoter methylation status, and calendar time. Results: Among 297 patients (244 pre-implementation, 53 post-implementation), baseline demographic and tumor characteristics were similar across cohorts. One-year overall survival was higher in the post-implementation cohort (60.4% vs 42.2%), corresponding to an adjusted hazard ratio of 0.61 (95% CI, 0.38-0.99). Postoperative MRI completion within 48 hours increased from 45.9% to 66.0% (adjusted cause-specific hazard ratio, 1.50; 95% CI, 1.05-2.14) with similar improvements observed at 7 days. Time to radiotherapy initiation did not differ between cohorts. Survival and MRI timeliness did not differ by rural or urban residence in either period, though rural radiotherapy delays were attenuated postimplementation. Conclusions: Systematic application of entrepreneurial methods to health system redesign was associated with clinically meaningful improvements in glioblastoma survival and care timeliness using minimal resources (single nurse navigator). These findings suggest that treating system change as an intervention grounded in user-defined priorities and oriented toward integrated systems rather than sequential process optimization can support sustainable transformation of complex, coordination-dependent care pathways and warrants evaluation in other disease settings.
Bai, L.; Liu, Y.; Tongye, H.
Show abstract
Background Glucagon-like peptide-1 receptor agonists (GLP-1RAs) are widely prescribed for type 2 diabetes and obesity, yet their neuropsychiatric safety profile remains incompletely characterized. We aimed to systematically evaluate neuro-adverse event (AE) signals for six GLP-1RAs and to validate key findings using population-based data. Methods We conducted disproportionality analysis of FAERS data for semaglutide, liraglutide, dulaglutide, tirzepatide, exenatide, and lixisenatide. RORs were calculated for 93 predefined neuro-AE MedDRA PTs across 11 neurological categories. External validation used NHANES 2013-2018 (n=17,057; 70 GLP-1RA users) with survey-weighted regression. Results We identified 41 significant neuro-AE signals. Semaglutide showed the strongest neuromuscular signal, muscle atrophy (ROR 3.94; 95%CI 3.42-4.54), corroborated by tirzepatide (ROR 2.35; 95%CI 2.04-2.71). Exenatide generated the highest psychiatric signal: nervousness (ROR 4.03; 95%CI 3.70-4.40). NHANES confirmed higher depression odds (OR 2.05; 95%CI 1.32-3.19; P=0.001) and reduced sleep hours (beta -0.35; P=0.033). Conclusions GLP-1RAs carry multiple neuropsychiatric safety signals, including muscle atrophy as a potential class effect and depression risk corroborated by population-level data. These findings support heightened clinical monitoring.
Glebov, K.; Zarovni, N.
Show abstract
The stall in clinical translation of extracellular vesicle (EV) therapeutics requires addressing critical and hidden causes of attrition and delays. Despite more than 150 registered clinical trials and fifteen years of clinical investigation, small EV (sEV) therapeutics have yielded zero FDA approvals, a translational deficit that has received remarkably little quantitative scrutiny. We systematically evaluated 783 EV-related clinical trials (152 of which therapeutic) registered through December 2025, benchmarking the sEV pipeline against 1,131 CAR-T cell therapy trials (ex-China) that share the "process is the product" manufacturing constraint, yet have delivered six FDA-approved products. A quality-failure paradox emerges, in which industry-sponsored sEV trials exhibit 39% higher composite methodological quality (quality index 2.49 vs. 1.79) yet fail at approximately five-fold the rate of academic programmes (28.6% vs. 5.1%; OR=7.40, 95% CI 2.5-22.3, p=0.0004). Cross-modality replication in CAR-T trials confirms the direction of this association (OR=2.51, p<0.001; Cochran-Mantel-Haenszel pooled OR=2.76, p<0.001). Registry abandonment, affecting 24% of academic sEV trials, constitutes a hidden failure mode that, when reclassified, dissolves the apparent paradox. Premature clinical entry of incompletely defined products, rather than insufficient methodological rigour, represents the central constraint on sEV translational progress. The opacity of research data, inadequate clinical evaluation, and failure to report negative (null) findings in clinical trials create a hidden crisis that drains hundreds of millions invested by public and private sector-eroding the return on investment-and traps progress in cycles of duplicated effort, wasteful resource allocation, and missed learning opportunities.
Li, Q.; Repalle, G. S. R.
Show abstract
Drug shortages represent persistent supply disruptions in the U.S. pharmaceutical market, threatening patient access and increasing drug costs. Prior research commonly treats shortages as binary events and relies on static designs, limiting insight into how shortage characteristics drive cost escalation. This study uncovers the heterogeneity behind drug shortages and pharmacy acquisition costs of generic non-injectable drugs. FDA drug shortage records with weekly National Average Drug Acquisition Cost (NADAC) prices were fit with fixed-effects models, duration-specific models, and a double machine-learning framework to characterize heterogeneity in shortage-price associations by duration, market structure, and shortage reasons. In the baseline two-way fixed-effects model, active shortage designation alone was not associated with a significant increase in NADAC under clustered standard errors. In duration-specific models, shortages lasting more than four consecutive weeks were associated with approximately 7% higher NADAC, while each additional cumulative shortage week was associated with approximately 0.37% higher NADAC. Estimated CATEs varied widely across drugs in each week. Allocation restrictions, raw material and distribution disruptions, together with a lack of manufacturers, characterized shortages with higher estimated CATEs. These findings support monitoring both shortage persistence and supply-chain mechanisms to mitigate impacts on healthcare systems.
Scherer, L. D.; Matlock, D. D.; Cronin, J.; Gritz, M.
Show abstract
Multi-Cancer Detection (MCD) tests can detect more than 50 different types of cancer using a blood test. Recently passed law in the U.S. guarantees that Medicare will pay for these tests when they are FDA approved and show evidence for clinical benefit. This manuscript provides estimates of the cost of MCD tests to Medicare under different assumptions of cost per test, eligibility, and screening uptake in the eligible population. This manuscript additionally estimates the cost of follow-up testing resulting from false positive results, which are considered avoidable costs caused by the screening test.
Farrow, E.; Balachandran, R.; Embleton, R.; Krogh, K.; Vollebregt, P. F.; Cornish, J.; Christensen, P.
Show abstract
Aims To develop the Bowel Irrigation Questionnaire (BIQ), a patient-reported experience measure (PREM) designed to assess the user experience of transanal irrigation (TAI). Methods Statements were generated through literature review and qualitative interviews with healthcare professionals (HCPs) and product users. Statements were rated on a 6-point content validity index scale through an international three-round online Delphi survey by 20 expert panel members. Consensus attainment was defined based on percentage agreement, statements which did not meet consensus were discussed at a final international online consensus meeting. The content validity of the PREM was evaluated through cognitive interviews and the Questionnaire on Questionnaires (QQ-10). Reliability was assessed using a test-retest design, where users completed the BIQ on two occasions one week apart. Results 215 statements were generated from 9 multi-disciplinary qualitative interviews and literature review. Statements were refined to reduce repetition and ensure clarity. 73 statements grouped into 11 domains were reviewed through the Delphi survey. Following the Delphi survey and clinical consensus meeting, the preliminary BIQ consisted of 15 items. Six cognitive interviews were conducted, resulting in a finalised BIQ of 16 items. 32 product users completed both the QQ-10 and test-retest study, the results of which demonstrated good content validity and temporal stability respectively. Conclusions The Bowel Irrigation Questionnaire is a novel PREM designed to assess the user experience of TAI in both clinical and research settings. The instrument demonstrates good validity, acceptability and temporal stability, supporting its use as a reliable measure of patient experience.
Dang, Z.; Ren, G.; Wang, Z.; Su, W.; Ma, Y.; Li, P.; Ji, D.; Li, L.; Gao, J.
Show abstract
Background: Under the DRG/DIP payment reform, the cost structure and its driving factors for laparoscopic cholecystectomy (LC) in resource-limited plateau regions remain unclear. Methods: Based on a single-center cohort of 605 plateau LC patients from May 2020 to October 2025, natural log transformation was applied to total hospitalization costs. Pearson/Spearman correlation, multivariate linear regression (traditional clinical model vs system-driven model with year dummies), and quantile regression were used. Results: Mean hospitalization cost 8097.49+/-936.85 CNY, CV=11.6%, Gini=0.062, demonstrating high homogenization. Traditional six-variable clinical model yielded R^2=0.008 (F=0.78, P=0.587), no significant predictors. The system-driven model achieved R^2=0.143 (F=3.42, P=0.001), with year dummies as dominant predictors. The study proposes the SAO (System-Allocation-Outcome) paradigm to replace the traditional SPO framework.
Cummins, J.; Drysdale, H.; Elson, M.; Hussey, I.; Goldacre, B.; DeVito, N. J.
Show abstract
Objective To evaluate the accuracy and cost of RegCheck, an automated large language model (LLM)-based workflow, for identifying clinical trial outcomes and detecting outcome misreporting by comparing its outputs with manual assessments from the COMPare Trials project. Design Validation study. Setting Sixty-two clinical trials originally assessed in the COMPare Trials project, sampled from five high impact general medical journals. Participants Published clinical trial reports and their corresponding prespecified registrations and/or protocols. Main outcome measures Four prespecified research questions were examined. RQ1 assessed outcome extraction recall relative to COMPare. RQ2 assessed accuracy of outcome classification as primary, secondary, or non-prespecified. RQ3 assessed accuracy of misreporting detection relative to COMPare, with additional manual adjudication of discrepancies between RegCheck and COMPare. RQ4 assessed the average per-paper cost of running the automated workflow. Results Across the validation papers, RegCheck achieved 91.2% outcome extraction recall relative to COMPare, and 83.6% outcome classification accuracy. For detection of outcome misreporting, RegCheck's overall accuracy was 85.6%. However, after resolving discrepancies with the original human judgements (which frequently favoured RegCheck's judgement), revised accuracy for outcome misreporting detection was 94.8%. The mean cost of running the workflow was 5.94 USD per paper. Conclusions RegCheck achieved high overall performance with a rigorous manual benchmark for identifying prespecified and reported outcomes in clinical trials, and detecting outcome misreporting, while operating at very low marginal cost. Adjudication of discrepant judgements suggested that RegCheck frequently identified valid issues not captured in the reference standard. Automated outcome checking may offer a scalable way to support editors, peer reviewers, and authors in detecting outcome switching and improving trial reporting.
Sohn, I.; Singh, T.; Carr, Z. J.
Show abstract
Background High-risk preoperative triage remains fragmented: existing tools often estimate risk without identifying modifiable mechanisms or linking classification to postoperative monitoring, destination planning, and rescue resources. This protocol describes implementation and evaluation of a Reserve-Stress-Rescue (RSR Framework), pathway that operationalizes perioperative high risk as a mismatch among patient physiologic reserve, procedural stress, and system rescue capacity. Approach RSR is a proposed clinician-facing, modular scoring framework for adults undergoing major surgery, especially patients with frailty, multimorbidity, poor functional capacity, anemia or malnutrition, cardiopulmonary disease, or limited postoperative support. Each domain, Reserve, Stress, and Rescue, is scored from 0 to 4 and recorded as both a three-part profile and a total score from 0 to 12. Scores map to Green, Amber, Red, and Crimson triage bands that trigger escalating actions, including targeted optimization, multidisciplinary review, anesthesia and surgical planning, postoperative destination selection, monitoring intensity, and predefined escalation criteria. Validation Plan The initial phase of this study received an exemption determination from the Yale University Institutional Review Board on June 3, 2026, under IRB Protocol ID 2000042729, with exempt categories 2(ii) and 4(iii), including a waiver of HIPAA authorization for access to and use of protected health information as described in the approved protocol. Evaluation will proceed in stages, assessing feasibility, interrater reliability, completeness, acceptability, discrimination, calibration, and clinical utility. Key outcomes include postoperative complications, unplanned escalation of care, intensive care utilization, failure to rescue, mortality, length of stay, triage burden, low-yield testing cascades, and management-changing pathway activation. Conclusion The RSR pathway reframes high-risk status as a modifiable interaction between vulnerability, operative insult, and rescue capacity rather than a fixed patient label. If feasible and valid, RSR may standardize high-risk identification, align perioperative resources with anticipated physiology, improve communication, and support safer, actionable shared decision-making.
Buianova, A. A.; Cheranev, V. V.; Kuznetsov, M. I.; Repinskaia, Z. A.; Belova, V. A.
Show abstract
Introduction: The application of pharmacogenomics (PGx) in pediatrics is limited by the lack of age-oriented interpretation approaches, as algorithms developed for adults do not account for ontogenetic changes in the activity of drug-metabolizing enzymes and transport proteins. The aim of this study was to evaluate the clinical applicability of pharmacogenomic data in Russian children, assess the concordance between genotype-based recommendations and the ontogenetic status of drug-metabolizing enzymes, and develop recommendations for the generation of age-oriented PGx reports. Methods: We analyzed whole-exome sequencing (WES) data from 524 pediatric patients and 635 newborns, filtering pharmacogenomic annotations according to PharmGKB/ClinPGx evidence levels (1A-2B) and the presence of the 'Pediatrics' tag. The concordance between genotype-based recommendations and the ontogenetic status of drug-metabolizing enzymes was assessed in newborns. In a pediatric subgroup of 100 patients, a retrospective analysis of medical records was performed to evaluate the structure of pharmacotherapy and the frequency of adverse drug reactions (ADRs). A 'PGx-ADR-cost' database was created, and the relative population burden index was calculated for 27 gene-variant-drug-ADR associations. Results: Clinically relevant annotations (requiring drug avoidance or dose modification) accounted for only 5% of all initial pharmacogenomic annotations in both cohorts; 67.6% (pediatric cohort) and 67.2% (neonatal cohort) of these were related to alleles with altered function. Concordance between genotype-based recommendations and the ontogenetic status of drug-metabolizing enzymes in newborns was observed in only 5 of 14 (35.71%) gene-drug pairs. ADRs were identified in 21% of the 100 pediatric patients; however, only two cases could be explained by high-evidence PharmGKB/ClinPGx annotations. Ranking by relative population burden identified UGT1A1*28-irinotecan-induced neutropenia and HLA-A*31:01-carbamazepine-induced severe cutaneous reactions as priority associations. Conclusions: Age represents a critical factor in the interpretation of pharmacogenomic data in children, as current approaches to PGx reporting do not adequately incorporate the ontogenetic context. We propose a pediatric PGx interpretation model that includes mandatory reporting of patient age, ontogenetic adjustment, evidence-level stratification, and multidisciplinary clinical assessment. Prospective validation is required to confirm the clinical utility of the proposed approach.
Farrow, E.; Coxon-Meggy, A.; Knight, L.; Bissett, I.; Bordeianou, L.; Boutros, M.; Burch, J.; Christensen, P.; Corrigan, N.; Croft, J.; Demian, M.; Dhadlie, S.; Emmertsen, K. J.; Gordon, K.; Sarah Faris-Sabboobeh, S.; Fearnhead, N.; Flavio FioreJr, J.; Keane, C.; Knowles, C.; Lloydwin, C.; Marinello, F.; Meggy, A.; Mohan, H.; Ng, K.-S.; Oliveira, C. L. P.; Oliveira, L.; Quyn, A.; Rose, A.; Stocken, D.; Warwick, A.; White, J.; Cornish, J.
Show abstract
Background The Low Anterior Resection Syndrome (LARS) score is an internationally validated instrument for identifying bowel dysfunction following anterior resection for rectal cancer. Although widely used, it has been shown to have limited sensitivity for capturing the impact of LARS on daily-life and response to treatment. We have therefore developed a novel patient-reported outcome measure (PROM): the LARS Impact and Consequences Assessment Tool (LARS-ICAT). Methods Initial development of LARS-ICAT followed a five-stage process following established PROM development guidance. Stage one established the conceptual foundation through previously published Delphi consensus. Stage two involved item generation, followed by evaluation of content validity through patient focus groups (n=11) in stage three. Stage four comprised iterative expert review and refinement through clinical consensus with patient involvement. Stage five involved cognitive interviews with patients conducted across five rounds (n=23). Results Several items identified through the Delphi consensus were reworded as they included multiple concepts. A one-month recall period was selected, with six and four response options for symptom and consequence items respectively. Additional consequence items, including impact on sleep and transport use, were incorporated. Focus groups and clinicians emphasised the importance of capturing individual symptom burden, leading to the addition of symptom bother scales. These iterative refinements culminated in LARS-ICAT v2.6. Conclusions LARS-ICAT is a novel PROM designed to assess symptom burden and treatment response in LARS. Future studies will assess its psychometric properties. Once validated, LARS-ICAT will provide a comprehensive, patient-centred assessment of LARS, enhancing our ability to manage this challenging condition.
Green, J. L.; Davies, H.; Russell, D. A.
Show abstract
Background: The relative merits of infrainguinal bypass and primary major lower limb amputation (MLLA) for chronic limb-threatening ischaemia (CLTI) remain uncertain, and the baseline profiles of patients selected for each strategy are poorly described. Methods: A systematic review and meta-analysis were undertaken in accordance with PRISMA 2020 and prospectively registered (PROSPERO: CRD42022356094). MEDLINE, Embase, CENTRAL, and CINAHL were searched from inception to March 2025. Prospective studies of adults with CLTI undergoing primary infrainguinal bypass or primary MLLA were eligible. Mortality, major adverse cardiovascular events (MACE) and subsequent amputation outcomes were synthesised using random-effects meta-analysis of proportions. Baseline comorbidity profiles were also extracted. Results: Twenty-seven studies involving 6,576 patients were included: 5,779 underwent infrainguinal bypass and 797 underwent MLLA. After bypass, pooled mortality was 3.7% at 30 days (95% CI 2.8%-4.9%, I2 = 49.4%), 18.5% at 1 year (95% CI 15.6%-21.9%, I2 = 62.3%), and 54.3% at 5 years (95% CI 50.5%-58.0%, I2 = 0%). After MLLA, pooled mortality was 9.2% at 30 days (95% CI 4.1%-19.3%, I2 = 73.5%), 28.5% at 1 year (95% CI 13.3%-51.0, I2 = 70.8%), and 39.9% at 2 years (95% CI 0.3%-99.3, I2 = 90.5%), although longer-term estimates were limited by sparse data and marked heterogeneity. Thirty-day MACE was 6.5% (95% CI 4.3%-9.7, I2 = 63.5%) after bypass and 2.8% after MLLA (95% CI 0.1%-37.6%, I2 = 0%). Early subsequent major amputation after bypass occurred in 3.9% of patients (95% CI 2.0%-7.7%, I2 = 91.2%), rising to 16.2% at 1 year (95% CI 12.6%-20.5%, I2 = 82.0%) and 33.3% at 3 years (95% CI 20.1%-49.8%, I2 = 0%). Early re-amputation after MLLA occurred in 10.9% of patients (95% CI 4.5%-24.4%, I2 = 40.3%). Baseline comorbidity burden was high in both groups, with substantial heterogeneity across studies. Conclusions: CLTI carries a poor prognosis regardless of treatment strategy. Infrainguinal bypass is associated with lower early mortality and better early limb preservation than primary MLLA, but long-term survival remains poor and later limb failure is common. Primary MLLA is not a low-risk alternative. Better contemporary comparative evidence utilising modern causal inference approaches is needed to support individualised decision-making.
Raisa, A.; Santaliz-Moreno, I.; Ayala, A.; Hamilton, J. G.; McQueen, A.; Souroullas, G. P.; Maki, J.; Waters, E. A.
Show abstract
Background: Epigenetics, the study of reversible changes in gene expression without altering the underlying DNA sequence, is increasingly applied in medical, commercial, and policy contexts. Yet, little is known about how this emerging science is communicated to the public. The purpose of this study was to examine communication strategies, sources, and modalities in epigenetic-related videos on YouTube- the most accessed platform for informal science education. Methods: We conducted a mixed-methods content analysis of 294 YouTube videos on epigenetics by conducting a keyword-based search on October 17, 2023. Video transcripts and meta-data were coded using a codebook developed both deductively and inductively. Qualitative analysis examined how communication strategies were used within videos and identified emergent themes (RQ1). Quantitative analyses examined the frequency of video and channel characteristics (RQ2), and presentation modalities (RQ3). Results: Findings reveal poor alignment with science communication best practices (RQ1): over 92% of videos failed to acknowledge scientific uncertainty, the comprehensibility level exceeded the recommended 8th-grade level (e.g., average readability grade 10.7), and professional research organizations were notably absent. Narrators were mostly male (56.7%) and white-presenting (73.7%) (RQ2). The majority of the videos used multi-modal strategies (e.g., visual texts mixed with animation and voice-over narration) to communicate epigenetic information (RQ3). Conclusion: Findings highlight the need for professional research organizations to be more proactive in public epigenetic communication efforts. Increasing narrator demographic diversity could broaden audience reach. Evidence-based communication tools are needed for health or science communicators discussing epigenetics on social media.
Zhang, Z.; Qadir, M. I.; Ramchand, R.; Belwadi, M.; Ball, R. P.; Konstantinopoulos, K.; Abbey, E. M.; Ernsberger, K. T.; Guzman, M. J.; Hendren, S.; Holcomb, B. K.; Robb, B. W.; Stankowski, T.; Waters, J. A.; Stefanidis, D.; Bilimoria, K. Y.; Mohanty, S.; Kolbinger, F. R.
Show abstract
Surgical video interpretation is a promising medical artificial intelligence application. However, no existing video annotation method preserves the spatiotemporal complexity of surgeon reasoning. Here we show that verbal reasoning and visual attention can be converted into structured, machine-actionable records of intraoperative behaviours. Our method decomposes transcribed verbal commentary into video-anchored semantic feedback chunks, which are classified via a large language model, with spatial grounding to surgical scenes via eyegaze or cursor tracking. We demonstrate method validity and scalability on structured and unstructured annotation tasks. For quality feedback on full-length colorectal procedures, the method reached near-human fidelity for chunking (mean cosine similarity: 0.95, SD: 0.01) and semantic classification across observations (mean Cohen's kappa: 0.71, SD: 0.07) and evaluative triggers (mean Cohen's kappa: 0.67, SD: 0.14), with excellent usability ratings. For structured critical view of safety assessment in laparoscopic cholecystectomy, implicit annotation yielded excellent agreement with explicit reviewer ratings (Cohen's kappa: 0.83, 0.49 and 0.81 across three criteria). We anticipate this method will advance surgical data science by enabling scalable construction of meaningfully annotated surgical video datasets.